Papers by Aashish Anantha Ramakrishnan
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)
Copied to clipboard
Ho Yin Sam Ng, Edward Hsu, Aashish Anantha Ramakrishnan, Branislav Kveton, Nedim Lipka, Franck Dernoncourt, Dongwon Lee, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ting-Hao Kenneth Huang
| Challenge: | Figure captions are crucial for helping readers understand and remember a figure’s key message. |
| Approach: | They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure . |
| Outcome: | The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles. |
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Text-to-image (T2I) models are trained on literal, object-centric prompts designed to reflect the visible contents of an image. |
| Approach: | They propose a method to extract key subjects and enhance their representation at embedding-level using Large Language Models. |
| Outcome: | The proposed model significantly improves image-caption consistency and human preference alignment. |
From Intentions to Techniques: A Comprehensive Taxonomy and Challenges in Text Watermarking for Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are rapidly growing and allowing textual content to be protected against unauthorized use. |
| Approach: | They present a unified overview of different perspectives behind designing watermarking techniques through a comprehensive survey of the research literature. |
| Outcome: | The proposed methods are based on the evaluation datasets used and watermarking addition and removal methods to construct a taxonomy. |
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships? (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on assessing factual and logical correctness in downstream tasks with limited emphasis on evaluating MLLMs’ ability to interpret pragmatic cues and intermodal relationships. |
| Approach: | They propose to use Coherence Relations to assess MLLMs' ability to perform multimodal discourse analysis using different prompting strategies. |
| Outcome: | The proposed model fails to match the performance of simple classifier-based benchmarks on 10+ MLLMs using different prompting strategies. |